Skip to content

feat(bench): v3 conditions, Wright Agent Score, and agent adapters - #464

Merged
Teakowa merged 16 commits into
mainfrom
feat/414-pi-devin-adapters
Oct 1, 2026
Merged

Teakowa merged 16 commits into
mainfrom
feat/414-pi-devin-adapters

Conversation

@e54-bot

@e54-bot e54-bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Refs #414, #466, #467. Brings the agent benchmark to the v3 contract: language-appropriate controls, a score card, and the adapters and wiki inputs it needs.

Conditions (v3). A cell is tool[+skills]/knowledge/network. tool is none, wright, or overpy (OverPy scenarios only); skills are wright-skill, workshop-skill, opy-skill, workshop-format-skill; knowledge is none, wiki (copied into the workspace), or web. Pairs not applicable to a scenario's language are skipped. The harness no longer adds text to the task prompt.

Contract wright-agent-bench/v3. status (completed, provider-interrupted, timeout, agent-error, invalid), usable with usableReason (lint errors, unsafe edits, and an unavailable grader block it), protocol, agentInfo, suite identity (version and hash), skill identities, harness commit, per-tool toolUse, and networkEnforcement disclosure. Canaries fail a run when a tool outside the condition is reachable.

Score wright-agent-score/v1. agent_bench.py score gives one card per language track on the canonical cell wright+wright-skill/none/off: scenario macro-average, two-stage bootstrap interval, Pass^k, exclusions, provisional reasons, and a refusal when runs come from different environments. report gains clustered intervals and --reference.

Scenarios. Held-out OPY and Workshop scenarios toward 8 per language, plus a negative for hallucinated names.

Adapters and wiki. pi, Devin, native Codex, and Antigravity adapters; a pinned wiki snapshot crawled by category and a local workshop-wiki skill build. Wiki content stays local and is not committed.

Verification. 64 unit tests pass with the pinned oracle installed (python3 -m unittest discover -s benchmarks/agent).

Not verified. Network off is declared, not enforced, and the card says so. No full-matrix run has been done on v3; the scenario count per language has not been checked against the 8 target.

…le canary

Refs #414. Adds adapters for pi and the Devin CLI that report per-turn usage, transcript, and loaded context, with host configuration kept out of the run (no context files or extensions for pi; an isolated HOME, no cross-tool rules, and denied MCP tools for Devin). Runs now fail a canary when instruction files exist above the workspace, the default output directory moves outside the repository, and results record whether network off was checked. Updates the matrix example and the benchmark contract.
Refs #414. wiki-snapshot fetches the Markdown mirror once into a local directory with per-document hashes and a snapshot hash. Runs with knowledge wiki require a snapshot and record its identity. Fetching goes through curl because the mirror returns 403 to Python's HTTP client.
Teakowa added 14 commits October 1, 2026 01:35
…ll and level

Refs #414. Local work, not pushed. wiki-snapshot now crawls category pages (the manifest lists only the first upstream page), with retries, resume, and bounded concurrency. wiki-skill builds a progressive-disclosure workshop-wiki skill from a snapshot, with OverPy spellings taken from the workshop-rs catalog and the opy-rs manifest and kept only when the pinned upstream compiler contains them. Adds the wiki-skill knowledge level, adapter support for a second skill, and line-buffered usage files so a killed run keeps its usage.
Reject changed snapshot and skill content before evaluation, correct the generated license notice, and align documentation and pilot matrices with the wiki-skill condition.

Refs #414
Run Codex and Antigravity with isolated configuration and normalized usage, isolate pi authentication, and restrict trial writes with the macOS sandbox. Keep unobserved context explicit and exclude provider failures from agent outcome metrics.

Refs #414
Clear stale pi provider errors after a successful response and classify observed transport failures as infrastructure exits. Refs #414.
Restrict outside-workspace instruction reads in the macOS file sandbox. Verify workspace instructions remain readable and host rules are denied before rerunning Devin pilot gaps. Refs #414.
Codex's workspace-write seatbelt cannot be applied inside the harness sandbox-exec profile (sandbox_apply: Operation not permitted), so every Codex tool write was denied in the preflight. Run Codex with its own sandbox off and let the harness file sandbox bound writes. Network off is declared-only for this adapter.
…guage

Refs #467. Adds greenfield-opy-elimination-race, understand-opy-events, modify-opy-kill-hud, modify-opy-extract-subroutine, and repair-opy-self-kill-score, plus Workshop twins of the last four (compiled from the OverPy references by the pinned compiler). All are test split, so each language track now has 8 held-out scenarios. Every scenario has a reference that passes, a seed that fails, and negatives that fail exactly the declared checks. Updates the ana-paintball hallucinated-name negative, which Wright now rejects like the upstream compiler.
…ght Agent Score

Refs #466, #467. Conditions become tool (none, wright, overpy) x skills (wright, workshop, opy, workshop-format) x knowledge x network, labelled tool[+skill...]/knowledge/network. The overpy tool and the language skills apply only to scenarios of their language, and matrix skips the rest. The harness adds no text to the scenario prompt. Results record protocol, Wright binary hash, skill hashes, suite hash, harness commit, agent info, and a status. usable now also blocks on unsafe edits and unavailable required graders. The wiki is copied into the workspace instead of linked. matrix stops after repeated provider interruptions and returns 3. New score command computes per-language scenario macro-average scores with a clustered bootstrap interval, Pass^k, and a score card. report pairs against any reference condition, splits by language, and uses clustered intervals.
@e54-bot e54-bot changed the title feat(bench): add pi and Devin adapters and an ancestor instruction-file canary feat(bench): v3 conditions, Wright Agent Score, and agent adapters Oct 1, 2026
@Teakowa
Teakowa merged commit 22fbdda into main Oct 1, 2026
15 checks passed
@Teakowa
Teakowa deleted the feat/414-pi-devin-adapters branch October 1, 2026 12:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

2 participants